Papers with audio representations

5 papers
Cross-Lingual Transfer Learning for Speech Translation (2025.naacl-short)

Copied to clipboard

Challenge: Increasing interest in building multilingual foundation models for NLP and speech research has led to limited data collection for training ST systems.
Approach: They propose to use Whisper to explore the behavior of multilingual speech foundation models with restricted data.
Outcome: The proposed model can translate to Chinese with a single language, and it can perform transcriptions in other languages.
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)

Copied to clipboard

Challenge: We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities.
Approach: They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations.
Outcome: The proposed model outperforms existing models on audio understanding tasks by 1%-84%.
Towards Cross-Lingual Audio Abuse Detection in Low-Resource Settings with Few-Shot Learning (2025.coling-main)

Copied to clipboard

Challenge: Online abusive content detection, particularly in low-resource settings, remains underexplored.
Approach: They propose to use pre-trained audio representations to detect abusive language in Indian languages using Few Shot Learning (FSL) .
Outcome: The proposed model can be used to classify abusive language in 10 languages using the ADIMA dataset with FSL.
Beyond Transcripts: A Renewed Perspective on Audio Chaptering (2026.acl-long)

Copied to clipboard

Challenge: despite its relevance, research on audio chaptering remains limited and predominantly textbased . authors: audio chapterers can't be used linearly because they skim, scrub timelines, jump to relevant moments . acoustic features and learning representations are not used for audio chapterer evaluation .
Approach: They propose to use audio-only architecture to automatically segment audio into coherent sections . they compare audio-based models with acoustic features and a novel audio-oriented architecture .
Outcome: The proposed audio-only architecture outperforms text-based approaches on acoustic features and LLMs.
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification (2025.naacl-long)

Copied to clipboard

Challenge: Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification.
Approach: They propose a training-free method that enhances audio and language representations using mutual feedback.
Outcome: The proposed method outperforms vanilla zero-shot evaluation with significant margins of 0.42%-27.0%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations